9-1 Beta Testing Plan — Worked Example (Godot Tetris)

How to use this: this is a completed plan at the scale your SAT expects — 3 testers, 3+ collection methods, a testing window of half a week plus a weekend. Read it beside C091-Beta-Testing-Plan-Template.md (same sections, filled in), then write the same kind of plan for your solution against your SRS. The 💡 boxes explain what each section is doing for the score — leave them out of your own plan.


Beta Testing Plan For:

Software Solution Name: Godot Tetris — a modern falling-blocks game (solo + two-player LAN versus), built in Godot 4.7

Target Users: students aged 12–18 who play casual games on school laptops; two players sharing a LAN for versus mode


1. Objective

Primary Goal: Confirm that the finished game is playable, understandable and stable in the hands of real players — the interface reads clearly, the modern mechanics (hold, ghost piece, hard drop, wall kicks) behave as players expect, and a two-player LAN match runs start-to-finish without failure.

Testing Focus:

  • Appearance (interface design, visual consistency)
  • Functionality (core features working correctly)
  • User Experience (ease of use, workflow efficiency)
  • Performance (speed, reliability, compatibility)

💡 The objective names the actual solution and what "working" means for it — not "test the game to find bugs". It targets appearance AND functionality AND requirements: that's the 3–4, 5–6 and 7–8 band behaviours in one sentence each.


2. Test Scenarios

Scenario 1 — Appearance Testing: "Can you read the board?"

Targets SRS non-functional requirements: usability (visual clarity), consistency of the interface.

Testers judge whether the play field, next-piece preview, hold slot, ghost piece, score and level display communicate without explanation.

Detailed User Steps:

  • Step 1 — Preparation: open the game on a school laptop; do not explain anything.
  • Step 2 — Starting Activity: tester starts a solo game from the menu unaided.
  • Step 3 — During Activity: after two minutes, tester points at each on-screen element and says what they think it shows (score, level, next piece, hold slot, ghost outline).
  • Step 4 — Ending Activity: tester plays until game over and describes what the game-over screen tells them.
  • Step 5 — Post-Activity Review: short survey items on visual clarity (1–5 scales).

Specific Feedback Points: Was the ghost piece understood without being told? Is the hold slot's "once per piece" state visible? Does the score/level area draw attention when a level-up happens?

Scenario 2 — Functionality Testing: "Do the mechanics do what players expect?"

Targets SRS functional requirements: rotation with wall kicks (FR-3), hold (FR-5), hard/soft drop and scoring (FR-7), line clears and levelling (FR-8).

Detailed User Steps:

  • Step 1 — Preparation: hand the tester the one-page controls card (←/→ move, ↓ soft drop, ↑/X and Z rotate, Space hard drop, C/Shift hold).
  • Step 2 — Starting Activity: tester attempts each control once in their first game.
  • Step 3 — During Activity: set tasks — rotate a piece flush against the wall (wall kick), hold a piece and retrieve it, clear two lines with one drop, reach level 2.
  • Step 4 — Ending Activity: observer checks the final score against the scoring table (100/300/500/800 × level, +1 per soft-drop cell, +2 per hard-drop cell).
  • Step 5 — Post-Activity Review: tester reports anything that "felt wrong" (a rotation that refused, a piece that locked too early, a hold that didn't respond).

Specific Feedback Points: Did any rotation near the wall surprise the tester? Did the 0.5 s lock delay feel fair or frustrating at speed? Did the displayed score match the events observed?

Scenario 3 — User Experience Testing: "A full versus match, cold"

Targets UX characteristics: learnability, efficiency, error tolerance. Targets SRS reliability requirement: a LAN session survives a complete match.

Detailed User Steps:

  • Step 1 — Preparation: two testers, two laptops, same LAN; neither has played versus mode.
  • Step 2 — Starting Activity: testers follow the lobby screen to connect to each other unaided (host + join).
  • Step 3 — During Activity: play one full match; observer logs any confusion, disconnection, or garbage-row event the players don't understand.
  • Step 4 — Ending Activity: the match ends; testers state who won and how they know.
  • Step 5 — Post-Activity Review: paired interview — what nearly stopped you, what would you change?

Specific Feedback Points: Time from "open game" to "match running" without help; whether incoming garbage rows read as an attack; whether the win/lose screen is unambiguous.

💡 Scenarios are real tasks with steps a stranger could run, not "click each button". Each names the FR/NFRs it exercises — that's the 7–8 band — and Scenario 3 targets UX characteristics by name, which is the 9–10 band behaviour.


3. Potential Users

User Who are they? Why selected? Available when?
User 1 Year 8 student, plays mobile puzzle games, has never played Tetris Reads the interface with fresh eyes — the learnability test can't be faked with an experienced player Lunchtimes this week
User 2 Year 11 student, experienced Tetris player (plays online guideline Tetris) Knows how hold, ghost and wall kicks should behave, so deviations from expectations surface immediately After school Thu/Fri
User 3 Parent, plays no games, uses a laptop daily for work Extreme-novice check on menus, controls card and game-over flow; also my weekend tester Saturday

Why these users represent your target audience: the target users are casual players on school laptops — Users 1 and 2 bracket that range (novice → expert), and User 3 tests whether the interface survives someone outside it. Users 1+2 together also form the LAN pair for Scenario 3.

💡 Each tester has a reason tied to what they reveal — that's "explains why potential users have been selected" (5–6 band). Role descriptions are fine; full names are not required.


4. Methodology

User Recruitment:

Ask in person this week; confirm each session time by message the day before. Consent forms handed out and signed before any session is recorded — no consent form, no session video.

Data Collection Methods:

  • Direct observation (watch users during testing)
  • Interviews (verbal feedback sessions)
  • Surveys/questionnaires (structured feedback forms)
  • Error logging (track problems encountered)

How Results Will Be Collected:

Method 1: Observation sheet (printed, one per session) — one row per event: time, what happened, tester reaction.

Method 2: Post-session survey (Google Form, 10 items: 1–5 scales on clarity/controls/fun + two open questions) — link sent as the session ends.

Method 3: Recorded interview (audio) using the six-question interview sheet; ~5 minutes per tester.

Method 4: Error log (spreadsheet) — every crash, disconnect, or "that's wrong" moment with steps to reproduce.

Data Validation Methods:

Comparison Method: survey answers cross-checked against what the observation sheet actually recorded — a tester who rates controls 5/5 but fumbled hold for three pieces gets a follow-up question in the interview.

Known Benchmarks: final scores checked against the scoring table; gravity/level progression checked against the level formula (level = 1 + lines ÷ 10).

External Validation: User 2's expectations from mainstream guideline Tetris act as the reference for "standard" mechanic behaviour.

💡 Four methods, and every method names its instrument — a survey that exists beats a "survey" that doesn't. Instruments must be ready before the first session: that's what the simulation checks.


5. Timeline

Phase Duration Activities
Setup ½ day Print observation sheets + controls cards, build the survey form, collect signed consent forms, test the LAN pairing on two school laptops
Active Testing 3 days (two weekdays + Saturday) Users 1 & 2 individually (Scenarios 1–2), Users 1+2 paired (Scenario 3), User 3 on Saturday (Scenarios 1–2)
Results & Feedback 1 day File raw data into the evidence folder, complete the error log, write the deviation note

💡 Total: half a week plus a weekend — the scale the assessment expects. Three testers, four sessions. A 6-week 12-tester plan (like the EasyRetail commercial exemplar) would fail the "realistic for the window" check.


6. Resources

Hardware/Software Requirements:

Two school laptops on the same LAN (versus needs both); the exported game build installed on each; any keyboard works — controls use physical key positions, not letters.

Support Materials for Users:

One-page controls card; consent form; the tester never sees the code or the plan.


7. Success Criteria

Primary Success Measures:

  • Every tester starts a solo game unaided in under 1 minute (learnability).
  • All set mechanic tasks (wall-kick rotation, hold, double line clear, reach level 2) completed by Users 1 and 2.
  • One full LAN versus match completes with no disconnection and an unambiguous result.
  • Displayed scores match the scoring table in every observed session.

User Satisfaction Targets:

  • Visual clarity and controls both average ≥ 4/5 on the survey.
  • No tester abandons a session.

Decision Framework:

What results would indicate success? All primary measures met and the error log holds only minor items → proceed to recommendations with priorities from the survey's open questions.

What results would require major changes? Any crash or LAN failure, a mechanic that confused both novice testers, or scores that don't match the table → these become the top recommended modifications in the C9-3/C9-4 report.

💡 Success criteria are countable afterwards — "≥ 4/5", "under 1 minute", "no disconnection" — so the report can say whether the test passed, not just how it felt.


Quality Checklist

C9-1 Assessment Criteria:

  • C9-1-1: Components clearly identified for testing
  • C9-1-3: Plan targets software appearance AND outlines potential users
  • C9-1-5: Plan targets functionality AND explains why users were selected
  • C9-1-7: Plan targets functional AND non-functional requirements AND documents how results will be collected
  • C9-1-9: Test scenarios target user experience characteristics AND documentation is clear and concise

Professional Standards:

  • All user selections include clear rationale ("why selected")
  • Multiple data collection methods documented ("how collected")
  • Test scenarios focus on relevant characteristics of your software solution
  • Timeline is realistic for user coordination and testing
  • All sections completed with clear, professional information